Missing value estimation for DNA microarray gene expression data: local least squares imputation

نویسندگان

  • Hyunsoo Kim
  • Gene H. Golub
  • Haesun Park
چکیده

MOTIVATION Gene expression data often contain missing expression values. Effective missing value estimation methods are needed since many algorithms for gene expression data analysis require a complete matrix of gene array values. In this paper, imputation methods based on the least squares formulation are proposed to estimate missing values in the gene expression data, which exploit local similarity structures in the data as well as least squares optimization process. RESULTS The proposed local least squares imputation method (LLSimpute) represents a target gene that has missing values as a linear combination of similar genes. The similar genes are chosen by k-nearest neighbors or k coherent genes that have large absolute values of Pearson correlation coefficients. Non-parametric missing values estimation method of LLSimpute are designed by introducing an automatic k-value estimator. In our experiments, the proposed LLSimpute method shows competitive results when compared with other imputation methods for missing value estimation on various datasets and percentages of missing values in the data. AVAILABILITY The software is available at http://www.cs.umn.edu/~hskim/tools.html CONTACT [email protected]

برای دانلود رایگان متن کامل این مقاله و بیش از 32 میلیون مقاله دیگر ابتدا ثبت نام کنید

ثبت نام

اگر عضو سایت هستید لطفا وارد حساب کاربری خود شوید

منابع مشابه

Missing Value Estimation for DNA Microarray Expression Data: Least Squares Imputation

Motivation: Gene expression microarray data sets often contain missing expression values. Robust missing value estimation methods are needed since many algorithms for gene expression analysis require a complete matrix of gene array values. In this paper, imputation methods based on the least squares and cluster structure are proposed to estimate missing values in the gene expression data, which...

متن کامل

An Improved Fixed Rank Approximation Algorithm for Missing Value Estimation for DNA Microarray Data

Gene expression data matriices often contain missing expression values. In this paper, we describe an improved fixed rank approximation algorithm (IFRAA) and compare it to the three recent methods for reconstructing missing entries for DNA microarray gene expression data: the Bayesian principal component analysis (BPCA), the fixed rank approximation algorithm (FRAA) and the local least squares ...

متن کامل

Weighted Local Least Squares Imputation Method for Missing Value Estimation

Missing values often exist in the data of gene expression microarray experiments. A number of methods such as the Row Average (RA) method, KNNimpute algorithm and SVDimpute algorithm have been proposed to estimate the missing values. Recently, Kim et al. proposed a Local Least Squares Imputation (LLSI) method for estimating the missing values. In this paper, we propose a Weighted Local Least Sq...

متن کامل

Collateral Missing Value Estimation: Robust Missing Value Estimation for Consequent Microarray Data Processing

Microarrays have unique ability to probe thousands of genes at a time that makes it a useful tool for variety of applications, ranging from diagnosis to drug discovery. However, data generated by microarrays often contains multiple missing gene expressions that affect the subsequent analysis, as most of the times these missing values are ignored. In this paper we have analyzed how accurate esti...

متن کامل

Evaluation of Missing Value Estimation for Microarray Data

Microarray gene expression data contains missing values (MVs). However, some methods for downstream analyses, including some prediction tools, require a complete expression data matrix. Current methods for estimating the MVs include sample mean and K-nearest neighbors (KNN). Whether the accuracy of estimation (imputation) methods depends on the actual gene expression has not been thoroughly inv...

متن کامل

ذخیره در منابع من


  با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید

عنوان ژورنال:
  • Bioinformatics

دوره 21 2  شماره 

صفحات  -

تاریخ انتشار 2005